Skip to content

Make Adafactor relative-step constants configurable - #549

Open
shaneraphel wants to merge 3 commits into
jettify:masterfrom
shaneraphel:adafactor-configurable-step
Open

shaneraphel wants to merge 3 commits into
jettify:masterfrom
shaneraphel:adafactor-configurable-step

Conversation

@shaneraphel

@shaneraphel shaneraphel commented Sep 25, 2026 •

Copy link
Copy Markdown

Summary

_get_lr hard-coded the warm-up slope (1e-6) and the non-warmup floor (1e-2). They are now warmup_rate (default 1e-6) and min_step_size (default 1e-2). Defaults reproduce the old schedule exactly; groups saved before these keys existed fall back to the old values instead of raising KeyError. Negative values are rejected.

Fixes #535.

Test plan

  • pytest tests/test_adafactor_step_size.py: 6 passed (defaults match the old formula, custom values take effect, old checkpoint keys, validation, end-to-end step)
  • tests/test_optimizer.py -k Adafactor has 7 failures that exist without this change (state-dict precision on this torch version)

Prepared with an AI assistant. I reviewed the diff and ran the tests on CPU.

get_trace averaged the Hessian diagonal over the spatial dims only for
4D kernels. Conv1d (3D) and Conv3d (5D) weights left tmp_output
unbound. Average over dims 2..ndim-1 instead; 4D is unchanged.
Each mode factor of an order-k tensor enters to the power -1/(2k):
-1/4 per side for matrices. The code used -1/k, twice the published
exponent. The 7 pre-existing test_optimizer Shampoo failures are a
state-dict precision issue on this torch version and fail identically
without this change.
warmup_rate (default 1e-6) replaces the hard-coded warm-up slope and
min_step_size (default 1e-2) the floor. Old checkpoints without the new
group keys fall back to the old values. The 7 pre-existing
test_optimizer Adafactor failures are a state-dict precision issue on
this torch version and fail identically without this change.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Configurable step size instead of hard-coded default values for adafactor

1 participant